Skip to content

feat(ui): LLM benchmark set say what it will actually run - #1183

Open
anhappdev wants to merge 7 commits into
fix/benchmark-set-option-persistencefrom
feat/benchmark-set-ui
Open

anhappdev wants to merge 7 commits into
fix/benchmark-set-option-persistencefrom
feat/benchmark-set-ui

Conversation

@anhappdev

@anhappdev anhappdev commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #1180; the diff here is only the commits on top of it.

Start screen: set card

  • The header states the outcome ("Runs 2 of 6 benchmarks") and lists those benchmarks, each with its icon, backend and delegate.
  • Options are always visible: checkboxes ("select one or more"), or radio buttons ("pick one") when the set has max_selected: 1. A set with nothing selected shows a warning.
  • The "Backends" disclosure sets backend and delegate per benchmark, and has an "Apply to all" control.

Results screen: set tile

  • The header shows "X of Y benchmarks ran", and benchmarks that didn't run read "Not run".

Fix

  • All six llm tasks have their own id, icon and info entry, so the LLM set's info button no longer crashes and the tasks sort correctly.

Code

  • The set card, loose card and result tile are state-free widgets driven by callbacks. BenchmarkConfigSection wires them to BenchmarkState.
  • Tests: task_coverage_test checks that every task in tasks.pbtxt has an id, info and icon. There are also widget tests for the cards, and SCREENSHOT_DIR=<dir> writes phone-size PNGs.

Notes

Screenshots: all 13 shipped benchmarks with the LLM set expanded, 390pt @3x. "Before" is #1180's head.

Before After
Start screen, before Start screen, after
Results, before Results, after

tasks.pbtxt ships six llm tasks; the frontend tables only knew two.

- BenchmarkId.allIds omitted llm-3b, llm-3b-instruct, llm-8b and
  llm-8b-instruct. BenchmarkStore sorts with allIds.indexOf, which returns -1
  for a missing id, so those four floated above every other benchmark.
- getLocalizedInfo had a case for llm-1b only and its default branch throws.
  showBenchInfoBottomSheet is the sole caller, and the benchmark set passes
  benchmarks[0] — one of the -1 sorted tasks — so the LLM set's info button
  threw 'unhandled task id'. Five of the six llm tasks were affected.
- The icon tables mapped the same two ids, so the rest fell back to the
  generic MLCommons logo.

Adds a test that reads assets/tasks.pbtxt and asserts every task in it has an
allIds entry, resolvable localized info and its own icon, so a task added to
the config fails here rather than under the user's finger.

BenchmarkId.llm and .llmInstruct are renamed to .llm1b and .llm1bInstruct now
that there are six; all three references were in this change.
A set collapses a cross product of options into fewer controls (#1095), but
the card only ever showed the tick count — "1/3 options selected" — and never
the benchmarks that count produces. One LLM parameter tick queues two
benchmarks through the hidden dataset option set, and nothing on screen said
so. That blind spot is where the stored-settings bug in #1180 lived.

Config card:

- The subtitle states the outcome, "Runs 2 of 6 benchmarks", and the card
  lists those benchmarks with the backend and delegate each will use.
- A set whose options are all off says so instead of reading "0/3".
- Options are chips carrying Option.name; the previous rows printed the raw
  option id, so the LLM set offered "1b", "3b", "8b".
- Options are always visible. They are the primary control for a set, and the
  old layout hid them behind a second disclosure on the same row as the gear.
- An option set with max_selected: 1 renders as a single choice and says
  "pick one". The proto has always had the bound; the UI never expressed it.
- Backend and delegate stay per benchmark, as #1164 made them, under one
  disclosure now labelled Backends. An "Apply to all" control writes one
  backend to every benchmark in the set that offers it and reads "Mixed" when
  they differ. A benchmark with a single backend shows plain text and no
  picker; the delegate picker is omitted when there is no real choice.
- The set icon and info sheet come from the set, via BenchmarkSetInfo, rather
  than from benchmarks[0].

Results card: the header states how many of the set's benchmarks ran, and the
ones that did not are marked "Not run" rather than given a result of "N/A".

The two cards are extracted into BenchmarkSetCard and BenchmarkSetResultTile,
free of BenchmarkState and driven by callbacks, which is what lets them be
rendered in a widget test. BackendChoice and the delegate dropdown move to
backend_choice.dart alongside a DelegateChoice widget.

unit_test/ui/benchmark_set_card_test.dart covers the states and, with
SCREENSHOT_DIR set, writes a PNG of each so the design can be reviewed without
building for a device.
Adds a screenshot test that captures the config screen at 390x844 with the
app's own theme, so the cards can be judged in context — density, how they sit
under the GO section, where the fold lands — rather than as isolated widgets.

BenchmarkState has a private constructor and late-initialised native
dependencies, so BenchmarkStartScreen itself cannot be pumped. The list is the
real BenchmarkConfigList holding the real cards; only the chrome around it is
reproduced, with the values copied from benchmark_start_screen.dart.

To make that possible the loose-benchmark row becomes BenchmarkLooseCard,
matching BenchmarkSetCard: free of BenchmarkState, driven by callbacks. The
list itself is BenchmarkConfigList, and BenchmarkConfigSection is now just the
adapter that wires it to state. The fixture gains a set-less benchmark so the
screen shows both shapes.
setSurfaceSize resizes the render surface but leaves MediaQuery reporting the
test default of 800x600. Anything measured from MediaQuery therefore laid out
for the wrong screen: the GO circle is MediaQuery.width * 0.32, so it came out
256pt across and 272pt tall inside a 390pt-wide phone, and the screenshots
made the start screen look far more cramped than it is.

Setting tester.view.physicalSize and devicePixelRatio sets both. On a real
390x844 phone the GO section is 140.8pt, 17% of the height, and the list gets
599.2pt — enough for both set cards and the top of the next one.
flutter_test sets debugDisableShadows, which paints every elevated widget as a
solid black outline. The GO circle came out with a thick black ring users never
see, which reads as a design change next to a screenshot of the current app.

The capture helpers turn shadows on, repaint the captured subtree, take the
image and restore the flag inside the test body, where the binding's invariant
check expects it.
@anhappdev
anhappdev requested a review from a team as a code owner September 26, 2026 12:45
@github-actions

Copy link
Copy Markdown

MLCommons CLA bot All contributors have signed the MLCommons CLA ✍️ ✅

@anhappdev
anhappdev marked this pull request as draft September 26, 2026 13:31
@freedomtan

Copy link
Copy Markdown
Contributor

@mohitmundhragithub and @AhmedTElthakeb please check this UI changes.

@anhappdev anhappdev changed the title feat: make a benchmark set say what it will actually run ui: LLM benchmark set say what it will actually run Oct 5, 2026
@anhappdev anhappdev changed the title ui: LLM benchmark set say what it will actually run feat(ui): LLM benchmark set say what it will actually run Oct 5, 2026
The results screen shows every benchmark with its own icon, but the set card
on the start screen listed the same benchmarks behind a plain dot, both in
what will run and in the backends panel. The icon replaces the dot in both
places; a benchmark the selection leaves out is dimmed in the backends panel,
the way the results tile dims one that did not run.
Users found the option chips less clear than the old checkboxes. Pills side by
side ("Offline | Online") read as a pick-one segmented control, and an
unselected chip showed no box, so nothing said it could be ticked.

Each option is now a checkbox with its label, or a radio button when the
option set is bounded to one choice. They keep the one-row layout and the 44pt
tap height, and the rule next to the heading reads "select one or more"
instead of "any".
@anhappdev

Copy link
Copy Markdown
Collaborator Author

@freedomtan @farook-edev The screenshots in the PR body are updated with your feedback:

  • Show each benchmark's icon on the set card rows
  • Use checkboxes for set options instead of chips

@sonarqubecloud

sonarqubecloud Bot commented Oct 6, 2026

Copy link
Copy Markdown

@anhappdev
anhappdev marked this pull request as ready for review October 7, 2026 12:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants